Papers by Christopher M Homan

6 papers
How Many Ratings per Item are Necessary for Reliable Significance Testing? (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for estimating model reliability are based on a few output responses per item.
Approach: They propose a method to determine whether an existing dataset has enough responses per item to assure reliable null hypothesis statistical testing.
Outcome: The proposed method can help researchers make better decisions about how to collect data for AI evaluation.
NSL-MT: Linguistically Informed Negative Samples for Efficient Machine Translation in African Low-Resource Languages (2026.findings-acl)

Copied to clipboard

Challenge: In low-resource settings, models encounter too few examples to reliably distinguish grammatical patterns from noise.
Approach: They propose a negative space learning machine translation (NSL-MT) method that augments limited parallel data with synthetically generated violations of the target language’s grammar and explicitly penalizes the model when it assigns high probability to these violations.
Outcome: The proposed method delivers 3-12% BLEU gains for well-performing models and 56-89% gains for models lacking decent initial support.
Subasa - Adapting Language Models for Low-resourced Offensive Language Detection in Sinhala (2025.naacl-srw)

Copied to clipboard

Challenge: A major challenge in the field of NLP are the disparities between high- and low-resource languages.
Approach: They propose fine-tuning strategies that have not been previously explored for Sinhala in the downstream task of offensive language detection.
Outcome: The proposed models outperform baseline models on the Sinhala offensive language detection task.
GAIfE: Using GenAI to Improve Literacy in Low-resourced Settings (2025.findings-naacl)

Copied to clipboard

Challenge: Illiteracy is a predictor of many negative social and personal outcomes in underresourced countries, where few books exist that are suitable for children to learn to read from.
Approach: They propose to use generative AI to create culturally-engaging materials for learning in mali's vehicular language Bambara by multiplying the content by 10 times . authors propose to apply bias-aware tools to reduce illiteracy and improve learning outcomes through native language education.
Outcome: The proposed toolchain and workflow can be adapted to address low literacy in mali using generative AI.
Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope Speech (2025.emnlp-main)

Copied to clipboard

Challenge: a new study examines how users interact with LGBTQ+ news content . a corpus of 1,419,047 comments on 3,161 YouTube news videos is used to analyze the content - both positive and negative - of cable news outlets.
Approach: They analyze how users interact with LGBTQ+ news content via a corpus of 1,419,047 comments on 3,161 YouTube news videos of major US cable news outlets.
Outcome: The proposed classifier detects positive (hope speech), negative, neutral, and irrelevant content.
Bayelemabaga: Creating Resources for Bambara NLP (2025.naacl-long)

Copied to clipboard

Challenge: a lack of well-structured multilingual datasets remains a challenge for machine translation in under-resource languages.
Approach: They propose to create a multilingual dataset for machine translation in the Bambara language, the vehicular language of Mali.
Outcome: The proposed dataset is the most extensive curated multilingual dataset for machine translation in the Bambara language, the vehicular language of Mali.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations